Meaning of Data
- डेटा व्याख्या — UGC NET/JRF & Assistant Professor Paper-1
Syllabus Coverage: Sources, acquisition and classification of data; quantitative and qualitative data; bar chart, histogram, pie chart, table chart, line chart and mapping; data interpretation; data and governance।
Data refers to raw facts, figures, observations, symbols or responses collected for analysis।
डेटा ऐसे कच्चे तथ्य, संख्याएँ, अवलोकन या प्रतिक्रियाएँ हैं जिनका विश्लेषण करके अर्थपूर्ण निष्कर्ष निकाला जाता है।
Data–Information–Knowledge
| Stage | Meaning | Example |
|---|---|---|
| Data / डेटा | Raw facts | 45, 60, 75 marks |
| Information / सूचना | Organised data | Average marks = 60 |
| Knowledge / ज्ञान | Interpreted information | Class needs improvement |
| Decision / निर्णय | Action based on knowledge | Remedial classes organised |
Example: एक survey में 500 students में से 300 ने online learning को useful माना। Raw data: 300 students; Percentage: 300/500 × 100 = 60%; Interpretation: Majority of students consider online learning useful.
Sources of Data
A. Primary Data — प्राथमिक डेटा
Data collected first-hand by the researcher for a specific purpose। शोधकर्ता द्वारा किसी विशेष अध्ययन के लिए पहली बार स्वयं एकत्र किया गया डेटा।
Methods: Observation, Interview, Questionnaire, Schedule, Experiment, Focus-group discussion, Tests and measurement, Field survey, Sensors or direct digital records
Example: Researcher स्वयं 200 college teachers से questionnaire भरवाता है। यह primary data है।
Advantages: Study-specific and relevant; methodology पर researcher का control; usually current and original।
Limitations: Time-consuming; costly; training and fieldwork required; non-response and interviewer bias possible।
B. Secondary Data — द्वितीयक डेटा
Data already collected by another person or institution and reused for the present study। ऐसा डेटा जो पहले किसी अन्य व्यक्ति या संस्था द्वारा एकत्र किया जा चुका हो।
Sources: Census reports, government publications, UGC/AISHE/NSS/RBI reports, books and journals, institutional records, research reports and theses, data repositories and websites, newspapers and databases
Example: AISHE report से university enrolment data लेना secondary data है।
Advantages: Economical and quickly available; large geographical or historical coverage; useful for comparison and trend analysis।
Limitations: Definitions may differ; data may be outdated; accuracy cannot always be controlled; data may not exactly fit the research objective।
Exam Trap: Primary और secondary data का distinction data के स्वरूप पर नहीं, बल्कि किसने और किस उद्देश्य से एकत्र किया इस पर निर्भर करता है। If a teacher uses data collected by their own university for a new research question, it may function as secondary data for that study।
Other Classifications by Source
| Classification | Meaning |
|---|---|
| Internal data | संस्था के भीतर से — attendance, results, finance |
| External data | संस्था के बाहर से — Census, UGC, market reports |
| Published data | Reports, journals, websites में उपलब्ध |
| Unpublished data | Private records, field notes, internal files |
Data Acquisition — डेटा का अर्जन
Data acquisition is the systematic process of obtaining data from relevant sources।
Census Method
Population की प्रत्येक unit से data collect किया जाता है। Example: किसी university के सभी 5,000 students का survey।
Sampling Method
Population की कुछ representative units से data लिया जाता है। Example: 5,000 students में से 500 students का representative sample।
| Census | Sample |
|---|---|
| सभी units शामिल | कुछ selected units |
| अधिक समय और लागत | अपेक्षाकृत कम समय और लागत |
| Sampling error नहीं | Sampling error possible |
| Small population में useful | Large population में useful |
Important Precautions
- Objective must be clear।
- Operational definitions तय हों।
- Representative sample चुना जाए।
- Instrument valid and reliable हो।
- Missing and duplicate values check किए जाएँ।
- Informed consent and privacy सुनिश्चित की जाए।
Classification of Data
Classification means arranging data into homogeneous groups according to common characteristics। वर्गीकरण का अर्थ समान विशेषताओं के आधार पर data को व्यवस्थित समूहों में बाँटना है।
A. Qualitative Classification
Quality or attribute के आधार पर classification। Examples: Gender, religion, marital status, teaching method।
B. Quantitative Classification
Numerical values के आधार पर classification। Examples: Age, income, marks, height।
C. Chronological Classification
Time के आधार पर arrangement। Example: Enrolment from 2022 to 2026।
D. Geographical Classification
Place or region के आधार पर। Example: State-wise literacy rate।
E. Cross-sectional Data
एक ही समय पर अनेक units का data। Example: 2026 में पाँच universities की enrolment।
F. Time-series Data
एक unit/variable का अलग-अलग समय का data। Example: 2021–2026 के बीच university enrolment।
Quantitative and Qualitative Data
Quantitative Data — मात्रात्मक डेटा
Numerical data that can be measured or counted।
1. Discrete Data — Countable values; usually whole numbers। Examples: Number of students, books, classrooms। (40 students possible; ordinarily 40.5 students नहीं।)
2. Continuous Data — Measurement से प्राप्त data जो एक interval में कोई भी value ले सकता है। Examples: Height, weight, time, temperature। (160.2 cm, 160.25 cm आदि values possible हैं।)
Qualitative Data — गुणात्मक डेटा
Non-numerical attributes, categories, perceptions or meanings। Examples: Learning style, gender category, attitude, classroom experience।
Important Distinction: अगर satisfaction को 1–5 scale में code किया गया है, numbers केवल categories को represent कर सकते हैं। इससे variable automatically continuous नहीं बन जाता।
Levels of Measurement
| Scale | Nature | Permitted comparison | Example |
|---|---|---|---|
| Nominal | केवल categories | Equal/not equal | Blood group, religion |
| Ordinal | Categories with rank | Greater/less | Rank, satisfaction level |
| Interval | Equal intervals; no true zero | Addition/subtraction | Celsius temperature, IQ |
| Ratio | Equal intervals and true zero | All arithmetic operations | Age, weight, income |
Exam Traps:
- Nominal scale में order नहीं होता।
- Ordinal scale में order होता है, पर ranks के बीच equal distance आवश्यक नहीं।
- Interval scale में zero arbitrary होता है।
- Ratio scale में absolute/true zero होता है।
- 40°C को 20°C से "twice as hot" नहीं कह सकते।
- 40 kg, 20 kg का twice है क्योंकि weight ratio scale है।
Tabulation of Data
Tabulation means presenting data systematically in rows and columns।
Parts of a Table
- Table number
- Title
- Caption/column headings
- Stub/row headings
- Body
- Unit of measurement
- Headnote
- Footnote
- Source
Good Table Characteristics
- Clear and self-explanatory title
- Consistent units
- Mutually exclusive categories
- Appropriate totals and subtotals
- Source and relevant notes
- No unnecessary information
Frequency Distribution
Frequency means the number of observations occurring in a category or class।
| Marks | Frequency |
|---|---|
| 0–10 | 4 |
| 10–20 | 6 |
| 20–30 | 10 |
Key Terms
Cumulative Frequency: Running total of frequencies। संचयी आवृत्ति किसी class तक की कुल frequency होती है।
Graphical Representation
A. Bar Chart — दण्ड आरेख
Used to compare discrete categories।
- Bars have equal width।
- Bars are separated by gaps।
- Height or length represents magnitude।
- Vertical or horizontal presentation possible।
Types: Simple bar chart, Multiple bar chart, Component/stacked bar chart, Percentage bar chart, Deviation bar chart
Suitable For: Department-wise enrolment, State-wise literacy, Comparison of male and female students
B. Histogram — आयतचित्र
Used mainly for a continuous frequency distribution।
- Rectangles touch one another।
- No gap between adjacent classes।
- Area of rectangle represents frequency।
- Class intervals are placed on X-axis।
| Bar Chart | Histogram |
|---|---|
| Discrete/categorical data | Continuous data |
| Bars separated by gaps | Rectangles touch |
| Order may be changed | Class order cannot be changed |
| Height represents value | Area represents frequency |
| Bar width generally equal | Width follows class interval |
Unequal Class Intervals: When class widths are unequal, use:
Unequal intervals में केवल raw frequency को height बनाना misleading हो सकता है।
C. Pie Chart — वृत्त आरेख
Shows how a total is divided into components।
Quick Conversions
| Percentage | Angle |
|---|---|
| 10% | 36° |
| 20% | 72° |
| 25% | 90° |
| 40% | 144° |
| 50% | 180° |
| 75% | 270° |
Example: Total students = 800; Science students = 200. Percentage = 200/800×100 = 25%; Angle = 25/100×360 = 90°
D. Line Chart — रेखा आरेख
Best suited to show change or trend over time। Examples: Year-wise enrolment, Monthly expenditure, Annual pass percentage, Population growth
Interpretation: Upward slope → increase; Downward slope → decrease; Horizontal line → no change; Steeper slope → faster change
Exam Trap: Steepness depends on the scale of axes। Truncated or unequal axes may exaggerate change।
E. Table Chart
Exact numerical values arranged in rows and columns।
Advantages: Precise comparison possible; suitable for ratio, percentage and average questions; multiple variables can be shown together।
Limitation: Trend is less immediately visible than in a line chart।
F. Mapping of Data
Spatial data is represented using maps।
| Map | Use |
|---|---|
| Choropleth map | Regions shaded according to rate/percentage |
| Dot map | Distribution or occurrence |
| Proportional-symbol map | Symbol size represents value |
| Flow map | Movement of people, goods or information |
| Cartogram | Region size modified according to data value |
Exam Trap: Population totals और population rates अलग हैं। Choropleth maps generally work better with rates, ratios or percentages, not raw population totals।
Which Representation Should Be Used?
| Purpose | Most appropriate representation |
|---|---|
| Category comparison | Bar chart |
| Continuous frequency distribution | Histogram |
| Part-to-whole relationship | Pie chart/component bar |
| Time trend | Line chart |
| Exact numerical values | Table |
| Regional pattern | Map |
| Relationship between two variables | Scatter plot |
Essential DI Formulae
Important: Percentage increase और percentage decrease के denominators अलग होते हैं। If value changes from 100 to 120: Increase = 20%। If it returns from 120 to 100: Decrease = 20/120×100 = 16.67%। Therefore, 20% increase followed by 20% decrease does not restore the original value।
Share from a Ratio: If male:female = 3:2 → Male Share = 3/5, Female Share = 2/5
Solved Example Set
The table presents students enrolled in five courses:
| Course | 2024 | 2025 |
|---|---|---|
| A | 200 | 250 |
| B | 300 | 360 |
| C | 250 | 300 |
| D | 400 | 440 |
| E | 350 | 350 |
2024 में total enrolment कितना था?
- A. 1,400
- B. 1,450
- C. 1,500
- D. 1,550
Explanation: 200+300+250+400+350 = 1500
Course A में 2024 से 2025 तक percentage increase कितना है?
- A. 20%
- B. 25%
- C. 40%
- D. 50%
(250−200)/200 × 100 = 25%
2025 में Course B और Course D के students का ratio क्या है?
- A. 9:10
- B. 9:11
- C. 10:11
- D. 11:9
360:440 = 9:11
किस course में enrolment में कोई परिवर्तन नहीं हुआ?
- A. Course A
- B. Course B
- C. Course D
- D. Course E
Explanation: दोनों वर्षों में enrolment 350 है।
2025 का average course enrolment कितना है?
- A. 330
- B. 340
- C. 350
- D. 360
(250+360+300+440+350)/5 = 1700/5 = 340
Data Interpretation Approach
Step 1: Read the Title
Data किस विषय, period और population से संबंधित है?
Step 2: Check Units
- Actual number
- Percentage
- Thousands/lakhs/crores
- Degrees
- Ratio
Step 3: Observe Base Value
Percentage किस total पर आधारित है?
Step 4: Identify Required Operation
Question difference, ratio, average, percentage या trend में से क्या पूछ रहा है?
Step 5: Estimate Before Calculation
Options widely separated हों तो approximation time बचा सकती है।
Step 6: Recheck Denominator
Percentage-related questions में denominator सबसे common error है।
Data Quality
Good interpretation requires good-quality data।
| Dimension | Meaning |
|---|---|
| Accuracy | वास्तविकता के निकट और error-free |
| Completeness | आवश्यक values missing न हों |
| Consistency | Different records में contradiction न हो |
| Timeliness | समय पर और updated |
| Validity | Defined format/rules का पालन |
| Uniqueness | Duplicate records न हों |
| Relevance | उद्देश्य के अनुकूल |
| Reliability | Repeated use में trustworthy |
Garbage In, Garbage Out: Incorrect or poor-quality input data produces unreliable conclusions।
Data Governance
Data governance is the framework of rules, roles, standards and accountability through which data is collected, stored, accessed, shared, protected and disposed of। डेटा गवर्नेंस नियमों, भूमिकाओं, standards और accountability की वह व्यवस्था है जिसके माध्यम से data का collection, storage, access, sharing, security तथा disposal नियंत्रित किया जाता है।
Core Elements
- Data ownership
- Data stewardship
- Data standards
- Metadata management
- Data quality
- Access control
- Privacy and security
- Interoperability
- Accountability and audit
- Retention and disposal
- Ethical data use
- Open-data policy
Data Governance vs Data Management
| Data Governance | Data Management |
|---|---|
| Policies तय करता है | Policies implement करता है |
| Who can do what? | How will it be done? |
| Authority and accountability | Operational processes |
| Standards and control | Storage, cleaning and processing |
Governance decides; management executes।
Data Life Cycle
Governance पूरे life cycle पर लागू होती है।
Open Government Data in India
The Open Government Data Platform India is designed, developed and hosted by NIC under MeitY। Its purpose includes facilitating access to government-owned shareable data in machine-readable form।
Open Data Principles: Availability, Accessibility, Machine readability, Reusability, Timeliness, Non-discrimination, Appropriate licensing, Privacy and security protection
Open Data ≠ Personal Data: हर government dataset को public नहीं किया जा सकता। Personal, confidential, security-sensitive और legally restricted data को protect करना आवश्यक है।
Personal Data Protection
The Digital Personal Data Protection Act, 2023 deals with processing digital personal data while recognising both personal-data protection and lawful processing needs। The official MeitY portal also lists the Digital Personal Data Protection Rules, 2025।
Exam-relevant Terms
- Data Principal: वह individual जिससे personal data संबंधित है।
- Data Fiduciary: वह entity जो processing का purpose और means तय करती है।
- Consent: Specific, informed and unambiguous permission।
- Purpose limitation: Data केवल declared lawful purpose के लिए use हो।
- Data minimisation: केवल आवश्यक data collect किया जाए।
- Security safeguards: Unauthorised access and breach से protection।
Data Ethics
Major Concerns
- Privacy violation
- Informed consent
- Surveillance
- Algorithmic bias
- Data manipulation
- Selective reporting
- Misleading visualisation
- Unauthorised sharing
- Re-identification
- Digital exclusion
Anonymisation vs Pseudonymisation
Anonymisation: Identity को इस प्रकार remove करना कि individual को reasonably identify न किया जा सके।
Pseudonymisation: Direct identifier को code से replace करना; additional information से identity पुनः connect हो सकती है।
Exam Trap: Pseudonymised data को हमेशा fully anonymous मानना गलत है।
Common Exam Traps
- Bar chart में gaps होते हैं; histogram में नहीं।
- Histogram continuous data के लिए है।
- Pie-chart sector angles का total 360° होता है।
- Percentage increase का base original value होता है।
- Average of percentages तभी सीधे निकाला जा सकता है जब denominators समान हों।
- Higher absolute value का अर्थ higher percentage आवश्यक नहीं।
- Correlation को causation नहीं माना जा सकता।
- Graph में truncated Y-axis difference को exaggerate कर सकती है।
- Missing data को automatically zero नहीं मानना चाहिए।
- Primary data हमेशा अधिक accurate हो — यह absolute statement गलत है।
- Secondary data हमेशा unreliable हो — गलत।
- Open data का अर्थ unrestricted disclosure of personal data नहीं है।
- Quantitative data numerical है, लेकिन हर numeric code ratio-scale data नहीं है।
- Unequal class intervals वाले histogram में frequency density आवश्यक हो सकती है।
- Data governance केवल cyber-security नहीं; इसमें quality, access, standards, ownership और accountability भी शामिल हैं।
Verified PYQs
Five organisations में कुल 35,000 employees थे:
| Organisation | Employees | Male:Female |
|---|---|---|
| A | 18% | 3:7 |
| B | 22% | 11:9 |
| C | 31% | 3:2 |
| D | 15% | 2:3 |
| E | 14% | 1:3 |
Question: सभी organisations में males की कुल संख्या कितनी है?
- A. 13,350
- B. 14,700
- C. 15,960
- D. 16,280
Explanation: A = 35000×18%×(3/10) = 1890; B = 35000×22%×(11/20) = 4235; C = 35000×31%×(3/5) = 6510; D = 35000×15%×(2/5) = 2100; E = 35000×14%×(1/4) = 1225। Total = 1890+4235+6510+2100+1225 = 15960
Five stores A–E sold 2,400 laptops। Their percentage shares were: A = 15%, B = 25%, C = 30%, D = 9%, E = 21%.
Question: यदि data को pie chart में represent किया जाए, तो Store E का central angle कितना होगा?
- A. 37.8°
- B. 75.6°
- C. 38.6°
- D. 77.2°
Explanation: Angle = 21% × 360° = (21/100)×360° = 75.6°। Total laptops की आवश्यकता नहीं क्योंकि percentage पहले से दिया है।
PYQ Trend Analysis
UGC NET Paper–1 में Data Interpretation सामान्यतः एक data set के आधार पर multiple questions के रूप में पूछी जाती है।
| Frequently tested area | Typical question |
|---|---|
| Table interpretation | Total, difference, ratio |
| Pie chart | Percentage, central angle |
| Bar chart | Year/category comparison |
| Line chart | Trend and percentage change |
| Combined data | Percentage + male:female ratio |
| Average | Category/year average |
| Ranking | Highest, lowest, second highest |
| Data sufficiency | Available data से answer संभव है या नहीं |
| Data governance | Privacy, access, open data, accountability |
| Misleading graphs | Scale, base and truncated axis |
Important Trend: Recent-style questions often require two or three operations: Total → Percentage Share → Ratio Share
इसलिए केवल formula याद करना पर्याप्त नहीं; सही sequence identify करना आवश्यक है।
December 2026 Expected MCQs
Continuous frequency distribution को represent करने के लिए सबसे appropriate graph कौन-सा है?
- A. Simple bar chart
- B. Pie chart
- C. Histogram
- D. Pictogram
Explanation: Histogram continuous class intervals को adjacent rectangles के माध्यम से दिखाता है।
एक category total का 35% है। Pie chart में उसका angle कितना होगा?
- A. 108°
- B. 120°
- C. 126°
- D. 135°
35/100 × 360 = 126°
किस scale में equal intervals होते हैं लेकिन true zero नहीं होता?
- A. Nominal
- B. Ordinal
- C. Interval
- D. Ratio
Explanation: Celsius temperature interval scale का उदाहरण है।
एक researcher Census report से literacy data लेता है। उसके अध्ययन में यह data होगा:
- A. Experimental data
- B. Primary data
- C. Secondary data
- D. Unclassified data
Explanation: Data researcher ने स्वयं पहली बार collect नहीं किया है।
एक value 200 से बढ़कर 250 हो जाती है। Percentage increase कितना है?
- A. 20%
- B. 25%
- C. 40%
- D. 50%
(250−200)/200 × 100 = 25%
दो groups की pass percentages क्रमशः 60% और 80% हैं। Combined pass percentage निकालने के लिए कौन-सी information आवश्यक है?
- A. केवल percentages
- B. प्रत्येक group के candidates की संख्या
- C. केवल total passed candidates
- D. Groups के नाम
Explanation: Different group sizes होने पर weighted calculation आवश्यक है।
Statement I: Histogram में adjacent rectangles सामान्यतः एक-दूसरे को touch करते हैं। Statement II: Bar chart केवल continuous data के लिए उपयोग किया जाता है।
- A. दोनों statements सही हैं
- B. दोनों statements गलत हैं
- C. Statement I सही, Statement II गलत
- D. Statement I गलत, Statement II सही
Explanation: Histogram continuous data के लिए है; bar chart categorical/discrete comparison के लिए।
Data governance का सर्वोत्तम वर्णन कौन-सा है?
- A. केवल data का computer में storage
- B. केवल cyber-attacks से protection
- C. Data-related policies, roles, standards and accountability
- D. केवल statistical analysis
Explanation: Governance data quality, ownership, access, privacy, security और accountability सभी को cover करती है।
किस data-quality dimension का संबंध duplicate records न होने से है?
- A. Timeliness
- B. Uniqueness
- C. Relevance
- D. Accessibility
एक graph का Y-axis 98 से शुरू होता है और 100 पर समाप्त होता है, जिससे छोटा difference बहुत बड़ा दिखाई देता है। यह किस समस्या का उदाहरण है?
- A. Sampling error
- B. Misleading scale
- C. Primary-data error
- D. Coding error
Explanation: Truncated axis visual difference को exaggerate कर सकती है।
Unequal class-width histogram में rectangle की उचित height किससे निर्धारित होगी?
- A. केवल frequency
- B. Cumulative frequency
- C. Frequency density
- D. Class midpoint
Frequency Density = f / Class Width
Assertion: Open government data transparency को support कर सकता है। Reason: सभी personal और confidential government records बिना restriction public किए जाने चाहिए।
- A. दोनों सही और Reason सही explanation है
- B. दोनों सही लेकिन Reason explanation नहीं है
- C. Assertion सही, Reason गलत
- D. Assertion गलत, Reason सही
Explanation: Open data transparency बढ़ाता है, लेकिन privacy, confidentiality और security restrictions लागू रहती हैं।
One-Page Revision Capsule
Classification
- Primary: first-hand
- Secondary: already collected
- Qualitative: attributes
- Quantitative: numbers
- Discrete: counted
- Continuous: measured
- Cross-sectional: many units, one time
- Time series: one variable across time
Graph Selection
- Category → Bar chart
- Continuous distribution → Histogram
- Part of total → Pie chart
- Time trend → Line chart
- Exact values → Table
- Regional distribution → Map
Formulae
Data Governance
Final Exam Approach
Join the Paper-1 Batch
Course fee FREE — ₹500 registration & academic/logistics contribution. Classes daily at 6:30 PM.